Meetings mostly in the morning.
Finalising setting up k8s server for monitoring.
Some discussion with Rizart. Was working on this today.
Remaking monitoring k8s cluster for database (up and running with PV).
However, do we need this now? Conversation w/ Rosie:
Rob Barnsley 4:14 PM
aside from the problem with rucio events not appearing in the dashboard, was there a particular reason why
we wanted to deploy our own database instance to capture these events?
4:16
if that problem is no longer presenting, and CERN are still planning to push rucio events to their internal
influx database, is there any good justification for pressing riccardo on this?
Rosie Bolton 4:16 PM
longevity of the database
4:17 I think the CERN stuff only holds 100 days?
Which was followed up with Riccardo/Rizart on rocket chat:
rbarnsley 4:25 PM
@rdimaria yes, thanks, it certainly looks a lot more promising (although I haven't quantitatively verified).
I believe the other justification behind migrating to hermes2 was to allow us to easily push events to an
externally hosted influx/es instance with a longer data retention policy
could someone remind me how long the events will live for on your infrastructure? we really need it for upwards
of 6 months.
ridona 4:31 PM
@rbarnsley For influx we have for sure 1 month (we could have also five years, there is such a retention policy,
I must check tho), for ES I need to ask, although my impression is its about 3 months or so
rbarnsley 4:26 PM
could someone remind me how long the events will live for on your infrastructure? we really need it for upwards
of 6 months.
rdimaria 4:32 PM
But in case we can surely ask CERN to extend these policies.
Amended Rucio dashboard for done/submission on transfer matrix.
Chat with Rizart/Rohini about what's required here. Waiting on Rizart to integrate Hermes2, then Riccardo to add to the Helm chart.
Will need to "redo" the cluster. One of the nodes is out of date.
Chat with Rizart in the morning.
Problem with rucio-client base container image and rucio-clients python package not authing.
Amended testing programme to include all sites.
Done. External access demonstrated.
Sci-Ops meeting.
Worked with Rohini to get influxdb externally accessible.
Tried to get python client to work w/ API.
Meeting IVOA, Rosie PDR, Escape DepOps, Sprint Planning
Removed duplicate bits from rucio-client with upstream. Still need to alter fts-analysis makefile and rucio-analysis makefile to run off this new base image.
Cleaning up a few bits, e.g. dockerhub integration with github for rucio-client-container.
Need to clean up container & remove bits from upstream Rucio.
Arranged demo with Philippa re: SDCSS on Thursday.
Meeting with Rizart re: events from Rucio. Going with hermes2 daemon, pushing events to a CERN hosted influxdb, and also SKA hosted influx/es.
Will need to add multiprocessing functionality & change workflow so it doesn't use a singer docker container?
Some problem with auth since 10:00. Hostnames have changed. Turned cronjob off for now.
Completed.
Not strictly req. this sprint (added story after the fact), but is required to demo to sci team.
SciOps meeting.
Simplified workflow, updated README and demoed to James.
Not strictly in ticket, but looked more at adding Rucio event transfer matrix. Got the impression that Rizart perhaps didn't know the scale of what was missing.
Sci-Ops meeting
Added a cronjob running this hourly on two sites: DESY-DCACHE and SARA-DCACHE.
Events on dashboard seem intermittent. Some turn up immediately, some turn up after a little time, and some don't turn up at all.
Have asked Rizart to investigate. He will get in touch with Riccardo.
Added a new, bigger volume for Shari (2Tb).
STFC user group meeting. DevOps meeting.
Amended code to do full meshed site test. Working on a Rucio level, but events not being passed through to dashboard.
Added bbcp to ASKAP instance and helped Shari push 300Gb file from Pawsey:
bbcp -z -P 10 -s 32 -w 2M -r image.restored.i.SB12835.cube_1665_sub.contsub.fits sbreen@130.246.212.43:/data
bbcp uses multiple sockets (-s) to transfer over TCP.
Added new instance with rucio-analysis toolkit. Tested and works.
Need to add cronjob for this, but what tests do we want and how often?
Created two new VMs to house keycloak and other services (API & Grafana). Works, but waiting for problem with RAL firewall (confirmed by Martin S.)
Remember need to switch the services running on the two instances, they're the wrong away around (so are the IPs!)
Switched over aeneas-ui to src-dmz. Waiting on testing from others.
Need to remember that username for RAL is now rob_barnsley_irisiam, not rbarnsley!
Created a new network for sdcss.
Meetings most of morning.
Rohini request aeneas-ui be redeployed with newer OS. Need to ensure users and ssh keys are ported (along with ip).
Openstack IPs now have "holes" in for necessary services.
Framework here:
https://github.com/ESCAPE-WP2/rucio-analysis
Replicates:
but is more configurable and dockerised. Should also be easily extensible to add other tests.
Meetings in afternoon.
Creating small python toolkit for "noise" tests.
Finished port to Minikube.
Emailed RAL about opening some ports for external testing.
Finishing porting over to Minikube.
PI8 planning.
PI8 planning.
PI8 planning.
PI8 planning.